Back

Statistical Methods in Medical Research

SAGE Publications

Preprints posted in the last 30 days, ranked by how well they match Statistical Methods in Medical Research's content profile, based on 11 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
Competing event regression on the relative subdistribution and cumulative-incidence scales

Mell, L. K.

2026-08-14 epidemiology 10.64898/2026.08.13.26360204 medRxiv
Top 0.1%
6.6%
Show abstract

In competing risks settings, covariate effects and group comparisons are usually assessed one event at a time - through log-rank or Cox tests on the cause-specific hazards, or Gray's test or Fine-Gray regression on a cumulative incidence function (CIF). This can obscure a clinically important quantity: the ratio between the event of interest and the competing event, since groups may differ little on the individual events yet differ sharply in their ratio. The generalized competing event (GCE) framework makes this ratio the object of inference; on the cause-specific scale the hazard ratio omega+(t) = lambda_1(t)/lambda_2(t) is estimated efficiently from a single stacked (Lunn-McNeil) model. We extend the framework to two scales that describe realized incidence. The subdistribution hazard ratio omega-tilde+(t) = lambda-tilde_1(t)/lambda-tilde_2(t) is estimated by a stacked, risk-set-weighted extension of the Lunn-McNeil construction; the cumulative-incidence ratio rho(t) = F_1(t)/F_2(t) - the odds that a subject's realized event by time t is the event of interest - by jackknife pseudo-observation regression of the Aalen-Johansen estimator. We relate the three contrasts: rho equals omega+ exactly under proportional cause-specific hazards, and equals omega-tilde+ only in the small-time limit under proportional subdistribution hazards, drifting toward 1 thereafter. The orthogonality that makes omega+ efficient is lost on both cumulative-incidence scales - omega tilde+ through overlapping weighted risk sets and shared censoring weights, rho through the shared all-cause survivor - so each carries a covariance term that must be handled and that bounds efficiency relative to the hazard-scale test. We derive the corresponding variances, study operating characteristics by simulation, illustrate on hypothetical prostate and head-and-neck cohorts, and provide an implementation in the gcemod R package.

2
Progesterone and hCG in expectant management success in tubal ectopic pregnancy: retrospective single-centre cohort study

Ahmad, A. K.; Pandrich, M.; Naik, A.; Astruc, A.; Lafferty, K.; Shah, N. M.; Ofili-Yebovi, D.

2026-08-07 obstetrics and gynecology 10.64898/2026.08.05.26359789 medRxiv
Top 0.1%
1.5%
Show abstract

Background: Early access to pregnancy assessment units now detects many tubal ectopic pregnancies (TEP) at a stage when they could resolve spontaneously, creating a management dilemma. Methods: We performed a hypothesis-generating exploratory analysis in a retrospective study to assess whether serum progesterone (P4) levels in women with TEP are associated with management outcome. Results: Ninety-one cases of TEP managed in a single centre over three years were analysed. Receiver operating characteristic (ROC) curve analysis was used to explore serum levels of progesterone (P4), first human chorionic gonadotropin (hCG) and peak hCG (alone and in combination) in relation with successful completion of expectant management. Decision-tree analysis using first hCG and P4 was additionally performed to explore clinical sequential risk stratification. 23% (n=21) successfully completed expectant management. P4 concentrations in the expectant management group (median 3 nmol/L, IQR 2.00 to 8.50) were significantly lower than in those requiring surgical or medical management (median 17 nmol/L, IQR 5.75 to 29.25; p=0.0002). Area under the ROC curve (AUC) values for P4, log10 first hCG, log10 peak hCG and P4 with log10 first hCG were 0.766, 0.814, 0.811 and 0.835, respectively, for predicting successful expectant management. However, hCG was not significantly outperformed. Nonetheless, Youden optimised thresholds for hCG and P4 are reported, alongside decision-tree analysis that identified sequential first hCG and P4 thresholds associated with successful expectant management. Conclusion: Lower P4 levels are associated with successful expectant management of TEP but they do not outperform hCG either alone or as an adjunctive marker.

3
Mitigating the Effects of Population Stratification in Gene-Gene Interaction Studies

Das, N.; Ueki, M.

2026-08-21 genomics 10.64898/2026.08.18.745398 medRxiv
Top 0.1%
1.4%
Show abstract

Population stratification is a major source of inflated false positive rates in genome wide association studies. However, relatively few studies have examined its impact on gene-gene interaction detection, despite the importance of epistasis for understanding the genetic architecture of complex traits. In this study, we identify scenarios under which population stratification can inflate the interaction test statistics. Through analytical derivations and simulation studies, we show that this inflation is not adequately controlled by including principal components as covariates in the regression model. We then propose an alternative approach that effectively controls the inflation of false-positive rates for interaction test statistics due to population stratification by using single nucleotide polymorphism-by-population structure interaction as an additional covariate term in the regression model.

4
Evaluating methodology to infer the effect direction in genetic association studies: applications to Body Mass Index, depression, and asthma

Chen, T.; Voorhies, K.; Reeson, A.; Seo, S.; Lee, S.; Hahn, G.; Hecker, J.; Prokopenko, D.; Hoth, K.; Kelly, R.; Lasky-Su, J. A.; Weiss, S.; Lange, C.; Lutz, S.

2026-08-10 epidemiology 10.64898/2026.08.05.26359780 medRxiv
Top 0.1%
1.4%
Show abstract

Mendelian Randomization (MR) is a popular tool for inferring causal relationships between traits using genetic variants as instrumental variables. These methods have been extended to also determine the direction of causality. However, causal direction cannot be inferred from a statistical test or estimation procedure (i.e. from data alone) without further assumptions and the methods operating characteristics and relative performances are not well understood. We conducted a comprehensive simulation study to illustrate this issue by evaluating type I error and power of 17 summary-based MR methods for inferring the effect direction. These methods fall within three methodological families: MR Steiger, Causal Direction (CD), and bidirectional MR approaches, with scenarios ranging across combinations of horizontal pleiotropy, unmeasured confounding, measurement error, longitudinal feedback, and varying sample sizes. While most methods achieved sufficient power levels under the alternative hypothesis in most scenarios, we found that every method was susceptible to inferring the wrong causal direction or under powered, and no method consistently maintained both correct type 1 error control and high power. In our applications, we evaluated the effect direction between the trait pairs body mass index (BMI) and major depressive disorder (MDD) and between BMI and asthma. To help researchers to evaluate the 17 methods to infer the effect direction and consider these challenges in their own data, we have developed MRdirection, an R package that runs the simulation studies examining the 17 directional MR methods across different user-defined scenarios. Our study, together with the accompanying R package, provides researchers with a tool for examining directional MR methods given different underlying assumptions.

5
Internal and External Validation of an Ensemble Learning Model Integrating Zygote Morphokinetics with Conventional Embryo Assessment for Blastocyst Prediction

ZHAO, M.; LIU, J.; HAN, D.; ZHANG, C.; ZHOU, Y.; CHEN, S.; LIU, C.

2026-08-23 obstetrics and gynecology 10.64898/2026.08.19.26359526 medRxiv
Top 0.1%
1.4%
Show abstract

Objective: To perform internal and external validation of a gradient-boosted decision tree (GBDT) fusion model that integrates zygote morphokinetic parameters with conventional embryo assessment features for blastocyst prediction, and to compare its discriminative performance against senior embryologists. Methods: This retrospective cohort study included 631 normally fertilized zygotes from 218 treatment cycles. A GBDT fusion model integrating 84 zygote morphokinetic parameters and 8 conventional assessment features was evaluated internally (5-fold cross-validation) and externally on a public dataset of 523 embryos with blastocyst outcomes. Model performance was assessed using area under the ROC curve (AUC), area under the precision-recall curve (AUPRC), F1 score, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV). Discrimination was compared with embryologist consensus using the DeLong test; agreement was assessed with Cohen's kappa. Results: The model achieved an internal AUC of 0.78 (95% CI 0.74-0.82), AUPRC 0.72, F1 0.73, sensitivity 0.74, specificity 0.77, PPV 0.72, and NPV 0.79. External validation on the public dataset demonstrated acceptable generalizability (AUC 0.76, 95% CI 0.71-0.81). The model significantly outperformed embryologist consensus (AUC 0.70, P<0.001) with moderate agreement (kappa=0.56). Decision curve analysis confirmed clinical net benefit at threshold probabilities of 0.15-0.55. Conclusions: The GBDT fusion model integrating zygote morphokinetics with conventional assessment demonstrates good discrimination and external generalizability for blastocyst prediction, providing an interpretable decision-support tool for embryo selection in IVF practice.

6
A Simulation Study Comparing Multiple Imputation and Complete Case Analysis for Handling Missing Preschool Body Mass Index

Savu, A.; Dover, D. C.; Hajihosseini, M.; Gaudet, L. A.; Kaul, P.

2026-08-14 epidemiology 10.64898/2026.08.13.26360115 medRxiv
Top 0.2%
0.8%
Show abstract

Background and Objective. Missing data frequently occurs in health databases and can bias analyses if not correctly dealt with. Using real-world data, we compared complete-case and multiple-imputation methods for recovering true parameters of a multivariable logistic regression model for the association between maternal glucose levels during pregnancy and child excess weight at preschool age, where missing values were present in as much as 30% of our sample. Methods. This study utilized a cohort of 130,424 children with complete preschool-age body mass index (BMI) measurements from the Calgary and Edmonton health regions of Alberta, Canada. In the complete BMI data, we introduced missingness through deletion following three distinct mechanisms: missing completely at random (MCAR), at random (MAR), and not at random (MNAR). To handle the missing data created, we employed complete-case and multiple-imputation methods. Maternal glucose levels during pregnancy were categorized into five groups and its association with child excess weight at pre-school age was determined based on a logistic regression model using the full observed data (yielding true values), observed data that was not deleted (complete-case estimates), and imputed data (multiple-imputation estimates). The accuracy of complete-case and multiple-imputation estimates were evaluated against the true values. Finally, we conducted a sensitivity analysis for the MNAR mechanism using pattern-mixture models with an additive shift. Results. Under MCAR and MAR, multiple-imputation generally outperformed complete-case, yielding smaller absolute and relative bias. Both methods achieved high significance ([&ge;] 0.96) for most effects. Mean squared errors for multiple-imputation and complete-case were similar missing completely at random, missing at random, and coverage was consistently high ([&ge;] 0.99). Under MNAR, both complete-case and multiple-imputation showed poor performance regarding bias and statistical significance. Sensitivity analysis using pattern-mixture models indicated performance varied by specific effect. Conclusions. Under MCAR and MAR, multiple-imputation introduced higher bias but demonstrated superior overall performance based on mean squared error and restored statistical power. Conversely, both methods failed under MNAR, where pattern-mixture modeling sensitivity analyses revealed highly variable, effect-specific performance due to unverifiable shift assumptions. When faced with missing data, researchers should assess missingness mechanisms, report both complete-case and multiple-imputation estimates under MCAR/MAR while accounting for power-versus-bias tradeoffs, and employ pattern-mixture sensitivity analyses to test robustness when MNAR is plausible.

7
Modelling the Effects of Smoking Behavior on Male-to-Male HPV Transmission and Anal Cancer Progression

Owolabi, R. O.; Martcheva, M.; Ghosh, I.

2026-08-12 epidemiology 10.64898/2026.08.11.26360159 medRxiv
Top 0.2%
0.6%
Show abstract

Human Papillomavirus (HPV) infection among men who have sex with men (MSM) has become a significant public health concern, particularly in countries where male vaccination is unavailable. Given the high susceptibility of MSM to HPV and anal cancer, and the unavailability of HPV vaccination for males in low- and middle-income countries (LMICs), there is a need to identify alternative interventions for reducing disease transmission and burden in this population. The novel mathematical model presented in this article couples smoking behavior dynamics with HPV transmission and anal cancer progression among MSM. Smoking reduction is introduced as an intervention to assess its effects on disease transmission and burden. The basic reproduction number (R0) is derived using the next-generation matrix method, and a global sensitivity analysis is performed using partial rank correlation coefficients (PRCC) to identify the influence of model parameters on RR0. Further, the theoretical analysis of the model reveals a backward bifurcation, implying that RR0 < 1 is necessary but not sufficient to eradicate the disease. The study finds that smoking reduction among MSM reduces HPV infection and anal cancer burden relative to baseline projections without intervention. The joint effect of smoking reduction and vaccination shows that the critical vaccination coverage needed to achieve RR0 <1 decreases as the level of smoking reduction increases. A similar outcome is observed for contact reduction. These findings highlight the importance of concurrent interventions, which can significantly curtail the spread of HPV and reduce disease burden in both the high-risk group and the general population.

8
Relevance Based Prediction: A Transparent, Non-Artificial Intelligence, Mathematical Solution to Personalized Opioid Treatment

Robinson, C. L.; Turkington, D.; Lee, L.; Kritzman, M.; Yong, R. J.

2026-08-10 pain medicine 10.64898/2026.08.07.26359966 medRxiv
Top 0.2%
0.5%
Show abstract

Accurate prediction of individual medical outcomes is essential for optimizing treatment allocation amid rising costs, coverage denials, and limited clinical resources. Traditional predictive models, including regression and neural networks, rely on average effects and cannot tailor predictions to the specific circumstances of individual cases. We present relevance-based prediction (RBP), a model-free method that predicts outcomes as weighted averages of observed cases, with weights determined by a rigorously defined measure of relevance. Unlike model-based methods that rely on fixed calibrated parameters, RBP revisits the original data for each prediction and customizes both the cases and variables used. Applied to opioid treatment, RBP provides case-specific insights unavailable from conventional models, including how each prior case informs a prediction, how each variable affects its reliability and value, and how reliable the prediction is before it is made. These individualized insights may prevent misleading average-based decisions and reduce harmful or suboptimal treatment.

9
Likelihood-Based Inference and Model Selection for Stochastic Gene Expression in Probability-Generating-Function Space

Wang, Y.; Shu, Z.; McAuley, K. B.; Cao, Z.

2026-08-25 systems biology 10.64898/2026.08.24.746673 medRxiv
Top 0.3%
0.5%
Show abstract

Selecting stochastic gene-expression models from single-cell counts requires accurate parameter inference and efficient model selection. Likelihood methods in count space can be costly when full stationary count distributions are unavailable, whereas approximate methods may lose accuracy. Probability generating functions (PGFs) offer a compact analytical alternative, but existing PGF workflows are generally not likelihood based and therefore rely on computationally intensive cross-validation. We develop a likelihood-based PGF framework for both tasks. Correlated empirical PGF values are used to construct a Gaussian quasi-likelihood for parameter inference and PGF-based Bayesian information criterion (BIC) for model selection. We show that the empirical PGF is exactly unbiased and that the parameter estimator is consistent, converges at the inverse-square-root sample-size rate, and is first-order asymptotically unbiased. For large samples and a uniquely preferred model, PGF-BIC selects the same model as leave-one-out cross-validation in PGF space.

10
Addressing Measurement Error of Machine-Learned Physical Activity in Nonlinear Dose-Response Survival Analysis: Development and Evaluation of Accelerated Failure Time, Spline, and Simulation-Extrapolation Method

Mamiya, H.; Zhang, Q.; Zhang, X.; Yan, Y.; Sharma, A.

2026-08-31 epidemiology 10.64898/2026.08.25.26361155 medRxiv
Top 0.3%
0.4%
Show abstract

Wearable (accelerometer) data and machine-learning allow objective assessment of the amount of daily physical activity. However, wearable-derived human activity is subject to measurement error. No studies have corrected the dose-response association between physical activity and survival time to chronic diseases, including cardiovascular disease (CVD). The objective is to estimate the measurement error-corrected association between CVD events and multiple measures of daily duration of light and total physical activity, derived from machine-learning and conventional accelerometer-processing methods. Our method combined an accelerated failure time model, spline, and simulation-extrapolation (SIMEX). The method recovered the true dose-response non-linear association in simulated data, while the naive model failed to capture it due to substantial attenuation. Application to the UK Biobank accelerometer cohort also showed an increased protective association of total physical activity after SIMEX correction (Time Ratio [TR] = 1.56, 95% CI: 1.28-1.82 vs. TR = 1.38, 95% CI: 1.24-1.54 for SIMEX-corrected vs. uncorrected dose-response association between the 95th and 5th percentiles of total activity), with a similar increase for light physical activity. Sensitivity analysis indicates that the female population experiences a substantially larger protective association after SIMEX correction than males. Dose-response survival analysis is a widely used analytical method in physical activity epidemiology and benefits from measurement error correction.

11
Bayesian Borrowing of External Information in Clinical Trials: A Comparison of MAP, RMAP, and SAM Priors

Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.

2026-08-31 pharmacology and therapeutics 10.64898/2026.08.26.26360843 medRxiv
Top 0.4%
0.4%
Show abstract

Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.

12
Describing health inequalities without distortion: Simple-Means MAIHDA vs Random-Effects MAIHDA

Merlo, J.; Bashir, N. Z.; Rodriguez-Lopez, M.; Khalaf, K.; Öberg, J.; Perez-Vicente, R.

2026-08-18 epidemiology 10.64898/2026.08.17.26360592 medRxiv
Top 0.4%
0.4%
Show abstract

Multilevel Analysis of Individual Heterogeneity and Discriminatory Accuracy (MAIHDA) describes health inequalities through three components: (i) specific contextual effects (SCE), (ii) general contextual effects (GCE), and (iii) discriminatory accuracy of the context. We present Simple-Means MAIHDA (S-MAIHDA), which estimates each stratum directly from its observed individuals, with no distributional assumption. The observed proportions are unbiased whatever the stratum size, and their confidence intervals report the uncertainty honestly. S-MAIHDA operationalises the three components on the probability scale. The SCE are the raw and standardised stratum prevalences and the modification of the sociodemographic average differences by the area. The GCE are the variance partition coefficient (VPC) and the contextual structuring of the between-stratum inequality, expressed as the contextual clustering of inequalities, the additive sociodemographic differences, and the contextual modification of inequalities (CMI). The contextual discriminatory accuracy is expressed by the area under the ROC curve (AUC), and the sensitivity and specificity at the population prevalence as the threshold for a possible intervention. Because its estimates are the observed data themselves, S-MAIHDA is the canonical description, and the compare diagnostic quantifies how Random-Effects MAIHDA (RE-MAIHDA), the usual implementation, departs from it: RE shrinkage pulls small strata towards the overall mean and can hide the very inequalities the analysis seeks. The approach is implemented in the smaihda Stata command and reproduced in free Python code. We illustrate S-MAIHDA on register data from Malmo, Sweden (43,291 individuals; 300 area-sociodemographic strata), showing how the three components separate two contrasting outcomes: psychotropic medication use, almost purely sociodemographic, stable across areas, with weak contextual structuring (VPC {approx} 4%, CMI {approx} 0%); and choice of a private general practitioner, strongly geographical (VPC {approx} 11%, CMI {approx} 17%), with the sociodemographic differences reshaped and amplified in wealthy areas. RE-MAIHDA attenuated inequalities. For describing inequalities, S-MAIHDA preserves what the data show.

13
Genetic Architecture and Sample Size Impact Relative Performance of Nonlinear Machine Learning and Standard Polygenic Risk Scores

Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.

2026-09-03 genetic and genomic medicine 10.64898/2026.08.29.26361109 medRxiv
Top 0.6%
0.3%
Show abstract

Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.

14
Primary Care Quality and Inappropriate Community Antibiotic Use: A Double Machine Learning Instrumental Variable Approach

Chen, Y.; Yi, H.; Rao, S.; Weber, A.; Hassmiller-Lich, K.; Sylvia, S.

2026-08-31 health economics 10.64898/2026.08.26.26361459 medRxiv
Top 0.6%
0.2%
Show abstract

Inappropriate antibiotic use presents a major global health challenge, particularly in low-resource settings where access to quality care is limited but antibiotics remain relatively unrestricted. This study estimates the causal effect of frontline primary care quality on inappropriate community antibiotic use, combining detailed community-based data from approximately 100 rural villages in rural China with an instrumental variable (IV) approach embedded within a double/debiased machine learning (DML) framework. We linked objective measures of village doctor clinical practice quality, measured through unannounced standardized patient visits, to household-level antibiotic use data collected from the same villages. To identify the causal effect, we constructed multiple candidate instruments from extensive provider characteristics and used an ensemble of machine learning algorithms within a flexible DML-IV framework to approximate an optimal instrument, addressing a many-weak-instruments problem. We found that improving village provider clinical practice quality reduced both antibiotic receipt during healthcare encounters for common diseases and household antibiotic storage for future self-medication. Our findings suggest that strengthening frontline primary care quality can meaningfully reduce inappropriate community antibiotic use without restricting access to essential treatment. More broadly, this study illustrates how causal machine learning can strengthen conventional causal estimation in complex observational settings in global health economics research.

15
Transformer-Based Survival Model for Cardiovascular Risk Prediction from Longitudinal Health Checkup Data

Tsurimoto, S.; Nomura, A.; Nagata, Y.; Noguchi, M.; Hirai, T.; Takeji, Y.; Tada, H.; Sakata, K.; Soichiro, U.; Okada, S.; Takamura, M.

2026-08-26 epidemiology 10.64898/2026.08.24.26361274 medRxiv
Top 0.6%
0.2%
Show abstract

Background: Cardiovascular disease (CVD) is a leading global health concern. Traditional models often miss nonlinear dependencies among physiological and behavioral factors. We hypothesized that a Transformer-based deep learning model, which excels at capturing complex patterns in structured data trained on large-scale health check-up records, would improve long-term CVD risk prediction. Methods: We analyzed longitudinal health records (2010?2024) from the Hokuriku Health Service Association (n = 100,056 without baseline CVD; development cohort). Incident CVD was defined as the first self-reported physician diagnosis of heart disease or stroke during the 10-year follow-up and was modeled as right-censored survival data. An external evaluation cohort comprised 79,756 Kanazawa City participants with health records. A Transformer model was trained using anthropometric, laboratory, and self-reported lifestyle data. Benchmarks included Cox regression, XGBoost survival embeddings, multilayer perceptron, the Framingham Risk Score, and the Hisayama Risk Score. Performance was evaluated using time-dependent area under the receiver operating characteristic curve (ROC-AUC) with a primary focus on the 10-year ROC-AUC, precision?recall AUC (PR-AUC), and concordance index (C-index). Interpretability was assessed through SHapley Additive exPlanations (SHAP) and a Feature-level Attention Network (FAN), visualizing the top 12 SHAP-ranked features to highlight key interactions. Results: In the development cohort, 4,113 CVD events (4.1%) occurred. The Transformer model achieved the best internal performance: 10-year ROC-AUC 0.821 (95% confidence interval [CI], 0.816?0.826), PR-AUC 0.427 (CI, 0.419?0.435), and C-index 0.781 (CI, 0.775?0.787). Performance remained robust externally (21,179 CVD events, 26.6%): ROC-AUC, 0.762; PR-AUC, 0.500; and C-index, 0.744. Regarding interpretability, SHAP identified age, electrocardiogram abnormality, antihypertensive medication, and sex as the most critical predictors. Notably, FAN elucidated the prognostic value of self-reported lifestyle factors. For example, daily exercise and weight gain modulated the model?s assessment of age-related risk. Within the attention network, age served as a central hub, linking these behavioral habits with physiological features. Conclusion: The Transformer-based model outperformed conventional methods in predicting long-term CVD risk. Model interpretation demonstrated the predictive utility of self-reported lifestyle factors, such as weight gain and daily exercise. These findings may support personalized CVD prevention and population-level risk stratification using routinely collected health checkup data.

16
Optimal LDCT screening for never-smoking Asian women using integrated polygenic and environmental risk: a microsimulation modelling study

Kowada, A.

2026-08-19 oncology 10.64898/2026.08.18.26360665 medRxiv
Top 0.6%
0.2%
Show abstract

Objective To identify optimal initiation ages and screening intervals for low-dose computed tomography (LDCT) screening among never-smoking Asian women using an integrated polygenic risk score (PRS)-environmental tobacco smoke (ETS) risk model, and to evaluate the cost-effectiveness of alternative screening strategies at these optimized ages. Design Integrated PRS-ETS microsimulation modelling. Setting Japan. Participants Never-smoking women stratified into eight risk groups defined by combinations of PRS levels and ETS exposure. Interventions LDCT screening at intervals of 1 to 10 years, annual chest radiography (CXR), or no screening. Main outcome measures Costs, quality-adjusted life years (QALYs), incremental cost-effectiveness ratios (ICERs), net monetary benefits, lung adenocarcinoma incidence and mortality, and optimal LDCT initiation ages. Sensitivity analyses used a willingness-to-pay threshold of US$50,000 per QALY gained. Results Optimal initiation ages ranged from 40 to 55 years across the eight PRS-ETS risk groups, with higher PRS-ETS risk associated with younger optimal initiation ages. Annual LDCT was the most cost-effective strategy across all PRS-ETS risk strata, yielding an ICER of US$40,471 per QALY in the lowest risk stratum and becoming cost-saving in higher risk strata. Over a lifetime, annual LDCT averted 8,534 lung adenocarcinoma deaths compared with annual CXR and 14,940 deaths compared with no screening. Conclusions Tailoring LDCT initiation age across integrated PRS-ETS risk groups maximizes mortality reduction achievable with cost-effective annual LDCT screening among never-smoking Asian women. These findings highlight an urgent limitation of global lung cancer screening guidelines that rely exclusively on smoking history and provide policy-ready evidence supporting the integration of PRS and ETS into future recommendations for precision LDCT screening for never-smoking populations.

17
A Principled Framework for Using Correlated Traits to Improve Risk Prediction

Akey, J. M.; Bierman, R.; Zhang, K.

2026-08-27 genomics 10.64898/2026.08.23.746504 medRxiv
Top 0.7%
0.2%
Show abstract

Although many complex phenotypes and diseases are influenced by shared genetic and environmental factors, risk prediction methods typically rely on genetic information from a single trait, leaving a rich source of predictive information largely unexploited. Phenotypic correlations can potentially be used to improve the accuracy of polygenic scores (PGS), but the conditions under which correlated traits meaningfully enhance prediction remain poorly understood. Here, we develop a general theoretical and simulation framework that quantifies the extent to which correlated "helper" traits improve predictive accuracy and identifies the factors that determine the magnitude of these gains. We show that helper traits can substantially improve predictive accuracy, with the magnitude of these gains governed by baseline model performance, genetic and environmental correlations, and the heritability of the target and helper traits, providing principled guidance for helper-trait selection. Paradoxically, when the target trait is itself weakly heritable, helper traits need not be highly heritable to substantially improve the accuracy of PGS, because low-heritability traits can still capture non-redundant environmental factors shared with the target trait. We empirically evaluated the use of helper traits by developing PGS models to predict type 2 diabetes using data from the UK Biobank. Helper traits substantially improved predictive accuracy relative to a single-trait PGS (AUC-ROC = 0.907 versus 0.677) and achieved performance comparable to models that use HbA1c (AUC-ROC = 0.889), the current clinical gold-standard biomarker. Our results establish a general theoretical and practical framework for exploiting correlated traits to improve polygenic prediction, provide principled guidance for selecting informative helper traits, and demonstrate how shared genetic and environmental architecture can be leveraged to substantially increase predictive accuracy. Furthermore, we developed an interactive web application to estimate the expected gain in accuracy from candidate helper traits using empirically measurable quantities.

18
Estimating age-specific heterogeneity in SARS-CoV-2 transmission from prospective longitudinal studies: the importance of correcting for study design

Chervet, S.; Layan, M.; Boëlle, P.-Y.; Guedj, J.; van der Werf, S.; Kerneis, S.; Sermet-Gaudelus, I.; Cauchemez, S.; Opatowski, L.

2026-08-10 epidemiology 10.64898/2026.08.06.26358866 medRxiv
Top 0.8%
0.2%
Show abstract

Longitudinal household studies, combined with mathematical modeling, are widely used to characterize the drivers of respiratory pathogen transmission, including the effects of age and symptoms. In practice, household recruitment protocols vary across studies, potentially introducing biases into observed data. However, these biases are typically overlooked in statistical inference, and their impact on parameter estimates remains unknown. Here, we use synthetic household outbreak data simulated under different recruitment protocols to evaluate how recruiting through infected children affects estimates of age-specific infectiousness and susceptibility. We show that, under child-based recruitment, the standard likelihood, which accounts only for transmission dynamics, leads to underestimating child infectiousness and overestimating child susceptibility by more than 30%. We then propose a novel estimation framework that explicitly incorporates the household recruitment process into the likelihood and show that it substantially reduces these biases. Applying this new approach to a French household study conducted during the COVID-19 pandemic, we estimated that children under 6 had 49% lower infectiousness than teenagers and adults during the Alpha wave, whereas no difference was observed during the Omicron wave. This study demonstrates that ignoring recruitment protocols can bias key epidemiological parameter estimates and highlights the importance of accounting for study design.

19
Product mix and time since cessation among Korean former smokers using non-combusted nicotine products: a KNHANES analysis with implications for lung cancer risk comparisons

Cook, S. F.; Cohen, G.; Cummings, K. M.

2026-08-13 oncology 10.64898/2026.08.11.26360188 medRxiv
Top 0.8%
0.1%
Show abstract

BackgroundObservational comparisons of former smokers who use non-combusted nicotine products with former smokers who quit without them require that two quantities be measured precisely: which product is being used, and how long ago cigarette smoking stopped. Neither quantity is recorded by the National Health Insurance Service (NHIS) screening instrument used in a recent Korean cohort study of post-cessation e-cigarette use and lung cancer risk. We characterized both quantities in a contemporaneous, nationally representative survey of the same population. MethodsWe analyzed the public-release microdata of the Korea National Health and Nutrition Examination Survey (KNHANES), 2018 to 2023, restricted to adults aged 19 years and older. Former smokers were identified by smoking status, and cessation duration was taken from the item recording months since the last cigarette. Former smokers currently using a heated tobacco product (HTP) or an e-cigarette (EC) were compared with former smokers using neither. KNHANES 2018 asked a generic e-cigarette question and, separately, a checklist naming HTP brands, allowing the two product classes to be separated. Distributions were compared with rank-based methods, the age-duration relationship with Theil-Sen regression, and residual imbalance by restricting the comparison group to respondents age-matched to within two years. ResultsThe 2018 analytic sample comprised 1,348 former smokers, of whom 43 currently used HTP or EC and 1,305 used neither. Among the product-using former smokers, 58% reported HTP use without e-cigarette use, 21% reported both, and 21% reported e-cigarette use without HTP use; 79% reported any HTP use. Median cessation duration was 0.7 years (IQR 0.25 to 1.5) among product users and 12.0 years (IQR 5.0 to 20.0) among those using neither (Kolmogorov- Smirnov D = 0.76, P < 0.001), with the product user having quit more recently in 92% of cross-group pairs. The separation persisted within the short-term (<5 year) stratum (D = 0.34, P < 0.001; 73% of pairs) and after age matching, where the residual gap was 9.3 years. Cessation duration rose with age among those using no product (Theil-Sen slope +0.30 years per year) but was flat among product users (-0.01). Restricting to the screening-eligible stratum used in the cohorts high-risk analysis did not attenuate the imbalance: among those aged 50 to 80, median cessation among no-product quitters rose to 15.5 years (n = 858), and adding a 20 pack-year criterion left 421 no-product quitters with a median of 11.0 years against three HTP/EC users who had quit 0.25, 1.0 and 2.0 years earlier, despite closely matched cumulative exposure (mean 37.6 vs 37.7 pack-years). The overall contrast reproduced in every wave from 2018 to 2023, with an age-matched residual of 9 to 11 years. ConclusionsIn a nationally representative survey of the same population and the same calendar year as the NHIS screening cohort analyzed by Kim et al., Korean former smokers using non-combusted nicotine products differed from other former smokers in two respects that bear directly on how such comparisons should be read. First, they were predominantly HTP users: 79% reported any HTP use, and only 21% reported e-cigarette use without HTP use. Second, they had stopped smoking approximately a decade more recently, a difference that survived stratification at five years and exact age matching. Neither quantity is recorded in the NHIS screening instrument. Cohort estimates comparing post-cessation product users with other quitters should therefore be interpreted with caution if they do not precisely characterize product composition and to time since cessation, and future studies should measure both directly.

20
Quantifying superspreading in bacterial STI outbreaks using phylodynamics

Sevilla, J.; Kende, J.; Duchene, S.; Meehan, M. T.

2026-08-17 epidemiology 10.64898/2026.08.14.26360404 medRxiv
Top 0.9%
0.1%
Show abstract

Bacterial sexually transmitted infections (STIs) pose a major global public health challenge, with Neisseria gonorrhoeae being of particular concern due to its persistently high prevalence and increasing antimicrobial resistance. The emergence of multidrug-resistant strains has narrowed treatment options, highlighting the importance of prevention. In this context, knowing whether there is superspreading (transmission heterogeneity) within a population becomes crucial for accurate public health measures. However, classic methods to quantify superspreading rely on dense contact tracing, and this is not always feasible. As an alternative, we can use Bayesian phylodynamic modelling to infer transmission dynamics, including superspreading. Yet modelling transmission dynamics using bacterial data remains problematic, although it is widely used for viral data. Here, we apply a multi-type birth-death model parametrised to quantify superspreading in N. gonorrhoeae outbreaks, estimating the fraction and relative impact of superspreaders and reproductive numbers for superspreaders and non-superspreaders. We also use a hierarchical modelling strategy with partial pooling to increase the power for detecting superspreading in each cluster. Model performance was successfully evaluated across a range of superspreading scenarios using both transmission-informed phylogenies and sequence data with phylogenetic uncertainty. Application to empirical genomic data revealed a substantial role of superspreading in N. gonorrhoeae transmission during the COVID-19 pandemic in Australia. These results highlight the impact of superspreading in N. gonorrhoeae transmission and the importance of detecting it to efficiently stop the dissemination of the disease